Artificial Intelligence & Machine Learning

NVIDIA Vera Rubin Arrives, Ushering in a Gigascale Era for AI Infrastructure

The artificial intelligence landscape is on the cusp of a profound transformation with the full-scale production ramp-up of NVIDIA’s Vera Rubin platform. This cutting-edge AI supercomputer, built from chip to grid, promises unprecedented performance, efficiency, and scalability, poised to redefine the capabilities of data centers worldwide. Early adoption and benchmark results from key partners like CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure underscore Vera Rubin’s potential to meet the exponentially growing demands of advanced AI workloads, particularly in the era of agentic systems.

NVIDIA’s strategic push with Vera Rubin signifies a monumental leap in its "AI Factory" strategy, characterized by a commitment to hyperscale, extreme codesign, and purpose-built infrastructure. The platform’s assembly involves a vast, mature rack-scale supply chain, spanning over 350 factory sites across 30 countries, a testament to the global effort required to meet the insatiable compute demand of AI development.

Advancing Performance Through Extreme Codesign

At the heart of the Vera Rubin platform lies an innovative approach known as "extreme codesign." This methodology moves beyond assembling disparate components; instead, it engineers seven specialized chips and five rack trays as a single, unified system. This integrated design philosophy encompasses the Vera Rubin NVL72, the Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX, and Vera BlueField-4 STX.

The NVIDIA Vera CPU, a cornerstone of this new architecture, is specifically engineered for the burgeoning "agent era" of AI. Its custom Olympus core delivers a remarkable 2x increase in single-threaded performance compared to previous generations, alongside a 3x boost in core-to-core bandwidth. Crucially, it achieves a 40% reduction in memory latency, positioning it as the most efficient single-threaded CPU for the increasingly complex agentic workloads that are becoming central to AI innovation. This focus on CPU efficiency is vital, as agentic systems can consume up to 15 times more tokens than traditional AI applications, placing significant strain on computational resources.

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

Accelerating AI Factories with Purpose-Built Networking

The platform’s networking capabilities are equally transformative. The sixth-generation NVLink scale-up technology provides more than double the throughput for complex workloads, a threefold reduction in latency, and a tenfold increase in packet rates when compared to standard Ethernet solutions. For scale-out networking, the Spectrum-X Ethernet system integrates 102.4T Spectrum-6 switch systems with 1.6T ConnectX-9 SuperNICs. This advanced suite features adaptive routing, sophisticated congestion control, comprehensive telemetry, and open operating system support, enabling 1.6 times higher RDMA bandwidth than conventional Ethernet.

Leading AI infrastructure builders, including CoreWeave, Microsoft, SpaceXAI, and Tesla, are among the first to integrate Spectrum-6 switches into their AI factories, signaling strong industry confidence in NVIDIA’s networking advancements. Furthermore, NVIDIA’s introduction of photonics with co-packaged optics for scale-out, the industry’s first such switch in volume manufacturing, offers a significant improvement in power efficiency, achieving 5x lower power consumption and a tenfold increase in Mean Time Between Failures (MTBI) compared to pluggable transceivers. Early adopters like CoreWeave, Lambda, and OCI are embracing this technology for its performance and reliability benefits. Spectrum-XGS Ethernet further extends this performance across distributed data centers, achieving 1.9x higher multi-site throughput, a critical factor for gigascale AI operations that are no longer confined to single locations.

NVLink Fusion, another key innovation, broadens the NVIDIA infrastructure platform’s compatibility by opening it to third-party XPUs (eXtreme Processing Units). This provides partners with a more expedited path to market, leveraging the established NVLink scale-up stack and its robust ecosystem.

Revolutionizing Efficiency: Setup Time, Water Conservation, and Cost Reduction

NVIDIA’s extensive experience in rack-scale codesign has culminated in a Vera Rubin NVL72 system that dramatically simplifies deployment. The absence of external cables, fans, or hoses within the compute tray slashes assembly time from hours to a mere minute, a substantial operational improvement for AI factory deployment.

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

The platform’s liquid cooling system operates with a 45-degree Celsius inlet temperature, enabling chiller-free dry-cooler operation. This higher-temperature dry cooling, coupled with the closed-loop liquid cooling, is projected to save millions of gallons of water per megawatt annually for new AI factories. This is a critical advancement in sustainability for the energy-intensive AI industry.

Beyond operational efficiencies, the economic benefits are equally compelling. Benchmarks, such as the one conducted by CoreWeave, indicate that Vera Rubin NVL72 delivers up to 10 times more tokens per megawatt and a tenfold reduction in cost per million tokens compared to the previous generation Grace Blackwell NVL72. This translates directly into more accessible and affordable AI development and deployment, democratizing access to advanced AI capabilities.

NVIDIA Vera Rubin Powers Europe’s Open Model Era with Microsoft and Mistral Partnership

In a significant development for European AI sovereignty, NVIDIA Vera Rubin is powering a newly expanded partnership between Microsoft and Mistral AI. This collaboration aims to bring frontier AI capabilities to the region, combining open European models with secure cloud and customer-controlled environments. The initiative is underpinned by a multibillion-dollar agreement focused on expanding AI infrastructure in Europe. Mistral AI is integrating thousands of the latest NVIDIA Vera Rubin GPUs to enhance AI compute availability, providing a unified platform for training, inference, and large-scale deployment.

This strategic alliance addresses Europe’s demand for advanced AI operating under its own legal frameworks and data control mandates. The Vera Rubin platform’s full-stack architecture, encompassing accelerated computing, networking, and software, ensures that models and agents can run efficiently from training through production. The partnership will deliver sovereign-ready AI across public cloud, cloud-connected, and fully disconnected private cloud environments. Mistral Medium 3.5 and OCR 4 models are now accessible through Microsoft Foundry, with Mistral models also integrated into Microsoft Copilot Studio. Azure Local and Foundry Local offerings will allow customers to utilize these models and tools across both cloud and customer-controlled environments. This empowers government agencies, healthcare and financial services organizations, and manufacturers to leverage AI for sensitive workflows and on-site data processing while maintaining stringent regional requirements for data control, governance, resilience, and strategic autonomy.

CoreWeave Benchmark Demonstrates Order-of-Magnitude Performance Leap

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

CoreWeave, a prominent AI cloud provider, has emerged as a leader in the early adoption of the Vera Rubin NVL72 platform. After extensive co-engineering efforts, CoreWeave became the first AI cloud to successfully bring up and validate the new system. Their independently conducted benchmarks on the DeepSeek-R1 model have revealed a staggering 10x improvement in tokens per second per megawatt when compared to the Grace Blackwell NVL72.

This metric, "tokens per megawatt," is considered the critical determinant for the profitable scaling of AI infrastructure. A higher tokens-per-megawatt ratio signifies greater AI intelligence derived from the same power budget or equivalent performance achieved with significantly less energy consumption. This efficiency gain is crucial for AI labs and enterprises, including those like Jane Street, who are reliant on scalable AI factories.

The DeepSeek R1 benchmark specifically highlights the Vera Rubin NVL72’s ability to overcome networking bottlenecks. Its 260 TB/s all-to-all NVLink 6 fabric effectively eliminates communication constraints, allowing the entire rack to function as a unified accelerator. CoreWeave’s deployment of NVIDIA Spectrum-X Ethernet SN6600-LD, built on the 102.4 Tb/s Spectrum-6 switch chip and featuring a liquid-cooled design, delivers unprecedented switching density and capacity, providing a fully non-blocking, multi-plane, multi-rail spine and leaf fabric that connects Vera Rubin NVL72 GPUs without oversubscription.

Google Cloud Launches A5X Instance Powered by Vera Rubin

Google Cloud has introduced its first A5X instance, powered by the NVIDIA Vera Rubin NVL72, to support the innovative work of London-based startup Ineffable Intelligence. Ineffable Intelligence is developing a new generation of "superlearner" systems that continuously learn through experience, aiming to drive breakthroughs across all scientific fields. Unlike traditional AI models that rely on static datasets, Ineffable’s agents learn directly from interactions with their environments through reinforcement learning. This process involves generating experience across massively parallel simulated environments and rapidly updating policies. These tightly coupled learning loops demand exceptional compute power, memory bandwidth, and interconnectivity, necessitating infrastructure capable of operating at immense scale with extremely low latency.

Lasse Espeholt, cofounder of Ineffable Intelligence, expressed his privilege in gaining early access to Vera Rubin, stating, "The next era of research requires the next era of hardware. We feel privileged to work with the teams at NVIDIA and Google Cloud, who were able to grant us early access to Vera Rubin. The support across both teams has been unmatched; we were up and running almost immediately and are already testing infra for our superlearners."

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

The NVIDIA Vera Rubin NVL72 is specifically designed for agentic training, offering predictable latency, high utilization, and significantly improved intelligence per dollar compared to previous systems. Google Cloud’s A5X instances, announced at Google Cloud Next, are bare-metal configurations built on NVIDIA Vera Rubin rack-scale systems. They promise up to a tenfold reduction in inference cost per token and a tenfold increase in token throughput per megawatt compared to the prior generation. These instances leverage NVIDIA ConnectX-9 SuperNICs and next-generation Google Virgo networking, enabling clusters that can scale to tens of thousands of NVIDIA Rubin GPUs within a single site and up to nearly a million GPUs across multisite configurations. This unified, AI-optimized stack is engineered to facilitate the training, tuning, and serving of frontier, open, agentic, and physical AI models, optimizing for performance, cost, and sustainability.

DeepInfra Benchmarks Highlight Vera CPU’s Orchestration Prowess

Independent benchmarks conducted by DeepInfra, an AI cloud platform and an early participant in the NVIDIA open AI ecosystem, demonstrate the significant performance advantages of the NVIDIA Vera CPU. DeepInfra’s tests reveal that the Vera CPU is more than twice as fast and can support a greater number of concurrent AI agents compared to alternative CPUs. The platform processes nearly 5 trillion tokens weekly, with approximately 30% driven by agentic systems, underscoring its focus on high-throughput AI inference.

The benchmarks indicate that the Vera CPU supports up to 1.6 times more concurrent AI agents at the same quality of service and achieves up to 2.2 times faster orchestration than competing CPUs, while simultaneously improving infrastructure utilization and cost efficiency. These findings underscore the Vera CPU’s capacity to deliver the cost efficiency, low latency, and throughput essential for production-level agentic AI. As AI agents tackle increasingly complex reasoning, planning, tool use, and data movement tasks, CPU performance has become a critical factor in orchestrating workloads around each model call. The Vera CPU’s design for agentic workloads, a key element of NVIDIA’s extreme codesign approach to AI factories, is proving instrumental in helping cloud providers enhance infrastructure utilization, boost cost efficiency, and support a larger number of concurrent AI agents with consistent quality of service.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button